Current Issue : October-December Volume : 2026 Issue Number : 4 Articles : 5 Articles
Speech is an accessible and information-rich clinical signal, but its diagnostic value is deeply entwined with biometric identity. Empirical and perceptual evidence shows that current anonymization methods cannot fully remove identity cues and often introduces disorder-dependent artifacts. These limitations reveal a need for a system-level approach. We propose a privacy stack that reconceives privacy as identity uncertainty, achieved through coordinated signal diversification, leakage-resistant model design, and privacy-aware infrastructure. This framework outlines a path toward clinically meaningful, equitable, and trustworthy speech-based artificial intelligence systems....
This systematic review investigates the integration of music education with science, technology, engineering, arts, and mathematics (STEAM) in South African basic schools. The study explores how music instruction can enhance scientific and technological literacy, particularly through acoustics concepts and ICT-enabled digital tools. Peer-reviewed journal articles, policy documents, and case studies published between 2000 and 2025 were systematically analysed following established review guidelines, with 19 studies meeting the inclusion criteria. The findings reveal that although music is formally embedded within the national curriculum, its interdisciplinary potential for STEAM integration remains underutilised. Key challenges include limited teacher professional development, inadequate ICT infrastructure, curriculum constraints, and insufficient emphasis on scientific principles such as sound waves, frequency, resonance, and vibration. However, innovative practices including digital composition software, virtual acoustics simulations, music-mathematics integration activities, and collaborative crossdisciplinary projects demonstrate positive effects on learner engagement, conceptual understanding, creativity, and problem-solving skills. The review further indicates that structured integration of digital music skills can support inclusive participation and foster critical thinking across artistic and scientific domains. The study concludes that embedding scientific and technological concepts within music education can contribute significantly to holistic learner development and interdisciplinary competence. These findings have implications for curriculum design, teacher training programmes, resource allocation, and education policy, supporting the advancement of contextually responsive STEAM education in South African primary and lower secondary schools....
This study takes the noise interference and sound quality attenuation problems of music signals during transmission as the core research, and constructs a noise reduction and sound quality optimization model for music transmission signals based on deep learning. The study collects typical music samples and multiple types of noise data, preprocesses the signals using time-frequency domain feature extraction methods, and on this basis, designs a deep neural network structure that integrates convolutional neural networks and perceptual optimization mechanisms. During the model training process, adaptive learning rates and perceptual loss functions are introduced to enhance the convergence speed and auditory consistency of the network in different noise environments. The experimental results show that this model outperforms the traditional spectral subtraction method, Wiener filtering, and ordinary CNN models in multiple indicators. Among them, the objective evaluation values such as PESQ, SDR, and STOI have significantly improved, the MOS subjective listening experience score reaches 3.87, and the naturalness of sound quality and the degree of detail restoration have been significantly improved. The model demonstrates excellent generalization ability and robustness under various noise types, effectively reducing background noise while maintaining the original timbre characteristics. The research results show that the combination of deep learning and perceptual optimization provides a new solution for noise reduction and sound quality enhancement of music signals, and can be widely applied in fields such as music production, audio restoration, and intelligent voice transmission. Future research will explore lightweight models and cross-modal optimization methods to achieve real-time high-fidelity audio processing....
The prosodic characteristics of a native language greatly influence early language acquisition. Yet, Japanese mothers are known to use a specific prosodic structure in infant-directed vocabulary (IDV)—specifically, three-mora, two-syllable words with a heavylight pattern—which, crucially, differs from the standard prosodic rhythm of adult vocabulary. This study used near-infrared spectroscopy to examine hemodynamic responses to the Japanese IDV form in 5-month-old ( n = 31) and 9-month-old ( n = 34) Japanese infants, targeting the period before and during the emergence of this preference. The results revealed that oxygenated hemoglobin was greater for the IDV form than for the non-IDV form in the left superior temporal gyrus (STG) for both age groups, consistent with the advantage of the IDV form observed in previous behavioral studies. Furthermore, this effect was localized to the left middle and left posterior STG in 5- and 9-month-old infants, respectively, highlighting early sensitivity to its prosodic structure followed by the emergence of phonological representation. This cortical shift, along with an observed trend toward adult-like patterns, may suggest a broader transition from perceptually accessible IDV structures to the more diverse patterns of standard adult vocabulary. Although 5-month-old infants who have not yet exhibited a preference for the IDV form may not have developed specific phonological representations, their brains’ ability to process its prosodic pattern could serve as a foundation for subsequent learning. These findings demonstrate that the specific structure of Japanese IDV acts as a foundational scaffold, guiding the transition from initial prosodic tuning to mature word-level processing....
This paper proposes a methodological framework for building a binary audio dataset for the automatic classification of fire sounds and forest ambience. Two operational recordings, one for the fire class and one for the forest class, are used strictly as seed data for controlled segmentation and augmentation. The workflow includes mono conversion at 16 kHz, amplitude normalization, segmentation into 5 s windows with 2 s overlap, lowintensity stochastic augmentation, and the systematic logging of the generated samples. The study also explains why augmented data are appropriate for training and internal validation, while final performance claims must remain reserved for testing on independent, standardized real recordings....
Loading....